‹ BackSSD POD

SSD POD

AI Inference
2026-07-23 09:42:12

AI inference shifts the memory stack as HBF, SSD POD and SRAM take on new roles

A TrendForce research note says the memory hierarchy behind AI systems is being redrawn as the industry shifts from training to inference. High Bandwidth Memory, or HBM, dominated the training era because it solved a bandwidth problem. In inference, the pressure point has changed. Key-value cache growth during prefill and decode is pushing capacity to the forefront, creating room for three different approaches: HBF for larger-capacity memory between HBM and SSD, SSD POD for offloading cache into lower-cost storage tiers, and SRAM for ultra-fast on-chip access in decode-heavy workloads. The report walks through how each technology fits into the stack rather than framing them as direct substitutes. HBF, being developed by SanDisk and SK hynix, is aimed at much larger capacity than HBM but has not reached mass production. NVIDIA’s SSD POD concept, tied to its Dynamo framework and NIXL transport library, is designed to move KV cache from limited GPU memory into CPU RAM, local SSDs, and remote storage without stopping inference. SRAM, used aggressively by companies such as Groq and Cerebras, trades capacity for speed by keeping memory on chip. The piece also points to recent product moves from NVIDIA, Google, SambaNova, Cerebras and Etched as evidence that inference-focused hardware design is accelerating.

1590
AI inference shifts the memory stack as HBF, SSD POD and SRAM take on new roles